Papers by Pedro Henrique Martins
Joint Learning of Named Entity Recognition and Entity Linking (P19-2)
Copied to clipboard
| Challenge: | Named entity recognition and entity linking are two fundamentally related tasks . most approaches focus on the mention detection part, assuming the correct mentions have been detected . |
| Approach: | They perform joint learning of named entity recognition and entity linking to leverage their relatedness. |
| Outcome: | The proposed model achieves competitive results with the state-of-the-art in both NER and EL tasks. |
∞-former: Infinite Memory Transformer (2022.acl-long)
Copied to clipboard
| Challenge: | Several efficient transformers have been proposed, but they all have a finite memory capacity and are forced to drop old information. |
| Approach: | They propose an unbounded long-term memory extension that extends the vanilla transformer by using a continuous-space attention mechanism to attend over the long-time memory. |
| Outcome: | The proposed model can model arbitrarily long contexts while keeping the computation budget fixed. |
Sparse Text Generation (2020.emnlp-main)
Copied to clipboard
| Challenge: | Current text generators require sampling from a modified softmax to avoid degenerate text . entmax sampling creates a mismatch between training and testing conditions . |
| Approach: | They propose to use entmax transformation to train and sample from a sparse language model to avoid degenerate text. |
| Outcome: | The proposed model improves fluency and consistency, fewer repetitions, and n-gram diversity closer to human text. |
Chunk-based Nearest Neighbor Machine Translation (2022.emnlp-main)
Copied to clipboard
| Challenge: | Semi-parametric models augment generation with retrieval, but require expensive retrieval operation for every generated token. |
| Approach: | They propose a semi-parametric model which augments generation with retrieval by retrieving tokens from a datastore. |
| Outcome: | The proposed model can retrieve chunks of tokens from the datastore, instead of a single token, with a low decoding speed. |